SEARCH RESULT

Year

Subject Area

Broadcast Area

Document Type

Language

1 results listed

2018 Diacritic Restoration of Turkish Tweets with word2vec

Social media platforms such as Twitter have grown at a tremendous pace in recent years and have become an important source of data providing information countless field. This situation was of interest to researchers and many studies on machine learning and natural language processing were conducted on social media data. However, the language used in social media contains a very high amount of noisy data than the formal writing language. In this article, we present a study on diacritic restoration which is one of the important difficulties of social media text normalization in order to reduce the noise problem. Diacritic is a set of marks used to change the sound values of letters and is used on many languages besides Turkish. We suggest a 3-step model for this study to overcome the top of the diacritic restoration problem. In the first stage, a candidate word producer produces possible word forms, in the second stage the language validator chooses the correct word forms and at the final word2vec is used to create vector representations of the words and make the most appropriate word choice by using cosine similarities. The proposed method was tested on both synthetic and real data sets, and we achieved a relative error reduction of 37.8% in our data sets compared to the previous study with an average of 94.5% performance.

International Conference on Advanced Technologies, Computer Engineering and Science
ICATCES

Zeynep Ozer İlyas özer Oğuz Findik

408 536
Subject Area: Computer Science Broadcast Area: International Type: Oral Paper Language: English